Assurance / Programme rubric icam-v1.3 judge claude · cortex
Programme assurance · ISO/IEC 25059

Programme assurance — Pilbara iron ore (pilot)

Quality of incident investigations across the corpus — not whether incidents occurred. Deterministic gates + an isolated ICAM judge score each closed investigation; flagged cases route to a human.
Investigations assured 412 closed & scored this quarter
Flagged for review 87 ≥1 dimension FAIL → HSE reviewer
Auto-assured rate 78% all dimensions PASS
Judge ↔ expert κ 0.74 ≥ 0.70 to gate · substantial
FAIL-reason distribution share of flagged investigations failing each ICAM layer · n = 87
Per-dimension verdicts aggregate into a systemic view — which aspect fails across many incidents, not one bad report.
Contributing-factor layer not substantively analysed
Organisational factors
73%
Task / environmental conditions
41%
Absent / failed defences
29%
Individual / team actions
9%
73% of site investigations miss the organisational-factors layer. A systemic gap — investigations stop at individual error and don't reach the latent, organisational "why." This is a programme finding, not a single deficient report.
Failing assurance dimension
Causal validity
61%
Completeness
48%
Consistency
44%
Factual grounding
12%
Compliance hard gate
0%
< 20% · isolated 20–60% · recurring > 60% · systemic
Measurable quality requirements ISO/IEC 25059 §11.1 · ICAM dimension ↦ characteristic ↦ target
Dimension ISO/IEC 25059 characteristic KPI target Actual Status
Completeness Functional completeness ≥ 90% PASS 91% on target
Causal validity Functional correctness ≥ 85% PASS 79% below target
Factual grounding Functional correctness · transparency ≥ 95% · 0 unsupported safety-critical claims 96% on target
Consistency Functional appropriateness ≥ 90% PASS 88% watch
Compliance Safety · freedom-from-risk 100% (hard gate) 100% pass
Targets turn "good rubric" into stated requirements with a target value (ISO/IEC 25030). Below-target dimensions are tuned with the client, not silently accepted.
Measurement-instrument validity gating
The LLM judge is itself a measurement instrument (§11.2), so its own reliability is a defined quality gate — measured against an expert-labelled gold set.
0.74 Cohen's κ
judge ↔ expert
0.0gate 0.701.0
Above the κ ≥ 0.70 required before the judge gates anything. Below that → rubric re-calibration, not deployment.
Per-dimension agreement · accuracy vs gold set
Completeness
0.89
Causal validity
0.71
Factual grounding
0.94
Consistency
0.82
Compliance
0.99
Causal validity is the weakest — the same dimension below its programme target. Re-calibration focus.
Re-validation
last validated2026-06-09 rubricicam-v1.3 locked judge modelclaude · cortex AI_COMPLETE gold set128 expert-labelled triggeron rubric / model change
Re-validated whenever the rubric_catalog content-hash or judge model changes. A standing human-audit sample detects drift and entrenched blind spots.
Societal / ethical risk (§11.4)
Operator-blame ratio
organisational-factor vs individual-action findings
1.6 : 1
Monitored so the rubric doesn't entrench systemic operator-blame bias. A ratio skewing toward individual actions would flag the judge, not the workforce.
Assist, don't decide
Every flagged case routes to a qualified HSE reviewer who decides assure / return-for-rework. The harness triages and explains with ICAM-layer + evidence rationale — it never closes an investigation. Safety-critical and regulated: human-in-the-loop is mandatory.